Papers by Harsha Vardhan Khurdula
Beyond Visual Understanding Introducing PARROT-360V for Vision Language Model Benchmarking (2025.coling-industry)
Copied to clipboard
| Challenge: | Current benchmarks for evaluating Vision Language Models (VLMs) often fail to thoroughly assess these models’ abilities to understand complex visual and textual content. |
| Approach: | They propose a benchmark that features 2487 visual puzzles designed to test VLMs on complex visual reasoning tasks. |
| Outcome: | The PARROT-360V Benchmark features 2487 visual puzzles designed to test VLMs on complex visual reasoning tasks. |